Journal of Biomedical Informatics
○ Elsevier BV
Preprints posted in the last 90 days, ranked by how well they match Journal of Biomedical Informatics's content profile, based on 47 papers previously published here. The average preprint has a 0.07% match score for this journal, so anything above that is already an above-average fit.
Wang, F.; Guo, Z.; Ye, Z.
Show abstract
Evidence-based medicine demands clinical answers that are not only fluent and medically plausible, but also anchored in traceable evidence, tailored to patient-specific clinical questions, sensitive to the hierarchy of evidence, and respectful of clinical safety boundaries. While general-purpose large language models (LLMs) exhibit strong medical language generation ability, they tend to lean on parametric memory, underuse retrieved evidence, hallucinate citations, conflate evidence levels, and draw conclusions that are not fully supported by the underlying literature. Such limitations pose particular risks in clinical decision support, where answer reliability, evidence traceability, and reasoning consistency are paramount. To address these issues, we present MedAgent, an evidence-based medical agent trained through an end-to-end pipeline that integrates supervised fine-tuning (SFT) cold start, reward modeling, and Group Relative Policy Optimization (GRPO). The agent is designed to execute a structured workflow encompassing clinical question understanding, PICO extraction, evidence retrieval, evidence stratification, citation-grounded answer generation, and quality evaluation. Specifically, a Qwen2.5-14B-Instruct backbone is first cold-started on 200 human-verified agent trajectories, equipping it with tool invocation, PICO parsing, structured response generation, and citation faithfulness. Next, a Qwen2.5-7B reward model is trained on 2{,}099 pairwise preference samples to provide semantic-level quality signals for evidence-based responses. Finally, GRPO reinforcement learning is conducted in a retrieval-augmented agent environment, where every rollout involves real evidence retrieval and is scored jointly by rule-based rewards and reward-model signals. To avoid over-reliance on training rewards, we further construct an independent evidence-based medical evaluation benchmark, MedTrustBench, which contains 200 clinical questions spanning 10 specialties and four difficulty levels. Each question is annotated with standardized PICO elements and rubric-based scoring criteria. The benchmark includes 1{,}187 rubrics across seven dimensions: question relevance, evidence hierarchy, evidence quality and timeliness, evidence-answer consistency, completeness and depth, logical rigor, and medical terminology. Under an identical RAG pipeline, retrieval tool, retrieval configuration, and evaluation protocol, MedAgentv17 attains 78.6 points, outperforming GPT-4.1 (75.3) and approaching GPT-5.4 (80.3). These results show that a 14B domain-aligned model can surpass strong general-purpose baselines on specialized evidence-based medical reasoning, while delivering practical advantages in cost, privacy, controllability, and hospital-oriented private deployment. The model and associated datasets are publicly released at https://www.modelscope.cn/profile/InfoxmedModel
Inyangala, J.; Mukudi, F. M.; Ojino, R.; Shisanya, M. S.
Show abstract
Background: Heart failure (HF) and chronic obstructive pulmonary disease (COPD) are among the leading causes of morbidity and mortality globally, with effective management heavily dependent on accurate severity staging using the New York Heart Association (NYHA) and Global Initiative for Chronic Obstructive Lung Disease (GOLD) classification systems. However, severity information is frequently embedded within unstructured clinical narratives rather than standardized Electronic Health Record (EHR) fields, limiting automated clinical decision support, disease surveillance, and retrospective healthcare analytics. Existing Natural Language Processing (NLP) approaches primarily rely on rule-based keyword extraction or supervised deep learning methods requiring large annotated corpora, which are often unavailable in many healthcare settings. Equally, most current systems inadequately integrate clinical ontologies for semantic reasoning and explainable classification, limiting interoperability and clinical applicability. Objective: This study aims to develop and evaluate an ontology-integrated NLP framework for automated extraction and severity staging of HF and COPD symptoms from de-identified clinical notes using NYHA and GOLD classification systems. Methods: The study will employ a Design Science Research (DSR) methodology to design, implement, and evaluate a hybrid NLP framework integrating rule-based extraction, SNOMED-CT ontology reasoning, and a Bidirectional Long Short-Term Memory with Conditional Random Field (Bi-LSTM-CRF) deep learning architecture for clinical sequence labeling. Approximately 1,000 de-identified clinical notes will be sampled proportionately from publicly available repositories including MIMIC-III/IV, eICU Collaborative Research Database, AmsterdamUMCdb, and MTSamples. Clinical text preprocessing will include tokenization, lemmatization, dependency parsing, abbreviation expansion, and negation detection. Ontology-guided semantic normalization will map extracted symptom entities to standardized SNOMED-CT concepts to support severity staging. Framework performance will be evaluated using precision, recall, F1-score, Cohens Kappa, sensitivity, specificity, positive predictive value, negative predictive value, confusion matrices, and correlation analyses against confirmed diagnoses and guideline-based severity classifications. Expected Outcomes: The proposed framework is expected to automate NYHA and GOLD severity staging across heterogeneous clinical note types without reliance on manually annotated severity labels. The ontology-integrated architecture is anticipated to improve semantic consistency, interpretability, and explainability of NLP outputs while enhancing EHR analytics, retrospective clinical audit, and AI-assisted clinical decision support. Conclusion: Findings from this study may provide a scalable and transferable framework for automated severity classification in data-rich but label-poor healthcare environments.
Malec, S. A.; Pradhan, M.; Upadhayaya, R.; Metzger, V.
Show abstract
Objective: Observational studies are essential for investigating risk factors for Alzheimer's disease and related dementias (ADRD), but inconsistent reporting and selection of covariates can contribute to residual confounding, omitted-variable bias, and reduced reproducibility. We developed and evaluated VAREX (Variable Extraction), a large language model (LLM)-based information extraction framework designed to automatically identify exposures, outcomes, and covariates from epidemiologic studies and populate structured evidence repositories. Materials and Methods: VAREX combines retrieval-augmented generation, biomedical language-model embeddings, semantic chunking, cross-encoder reranking, and prompt-engineered LLM workflows to extract epidemiologic variables from full-text biomedical articles. The framework was evaluated using a reference-standard corpus of observational studies examining blood pressure variability (BPV) and Alzheimer's disease-related dementias (ADRD), together with external validation datasets involving other exposure-outcome relationships. Extracted variables were compared with independently curated human reference standards using semantic matching and one-to-one assignment procedures. Covariates were additionally classified into ten epidemiologically relevant semantic categories. Results: In the primary BPV[->]ADRD corpus (10 studies), VAREX achieved a precision of 0.91, recall of 0.84, and F1-score of 0.87 for variable extraction. Covariate classification accuracy was 0.90, yielding a strict extraction-and-classification F1-score of 0.78. External validation datasets demonstrated comparable performance across diverse epidemiologic domains, with extraction F1-scores ranging from 0.73 to 0.85. Category-level performance was strongest for health behaviors (F1=0.96), sociodemographic variables (F1=0.90), and medication exposures (F1=0.89). Compared with published estimates of manual systematic-review effort, VAREX reduced processing time from approximately 61 minutes to 9 minutes per article, representing an 85.7% reduction in review time. Discussion: These findings demonstrate that LLM-based information extraction can accurately identify and classify epidemiologic variables across heterogeneous observational-study designs. Automated extraction enables scalable construction of structured repositories of exposures, outcomes, and covariates while substantially reducing the labor required for evidence synthesis and systematic reviews. Conclusion: VAREX provides an effective framework for automated extraction and classification of epidemiologic variables from the biomedical literature. By supporting large-scale evidence synthesis and structured knowledge resource development, VAREX may facilitate more rigorous observational research, improved confounder identification, and enhanced reproducibility in epidemiology.
Chinthala, L. K.; Lemon, C.; Shaban-Nejad, A.; Farage, G.; Davis, R. L.; Xu, H.; Madlock-Brown, C.
Show abstract
Objectives: This study aimed to leverage FLAN-T5-Large, BERT, RoBERTa, and Gemma-2-2B, with fine-tuning, to identify instances of social isolation and social support within unstructured clinical notes. Materials and Methods: Annotated clinical note spans containing social context cues were used to fine-tune each model. Performance was evaluated using Accuracy, Precision, Recall, and Macro-F1 score. A structured prompt was used to instruct the model to perform classification task and mitigate overgeneralization. Performance comparisons across the models assessed sensitivity, robustness, and false positive reduction. Results: FLAN-T5-Large achieved highest performance, with Macro-F1 of 0.92{+/-}0.04, demonstrating balanced results across classes: social isolation (F1 = 0.91{+/-}0.03), no social isolation (F1 = 0.94{+/-}0.05), and social support (F1 = 0.90{+/-}0.04). Gemma-2-2B produced comparable results, with Macro-F1 score of 0.89{+/-}0.10. BERT and RoBERTa achieved lower Macro-F1 scores of 0.77{+/-}0.17 and 0.80{+/-}0.21 respectively, with variability across categories. Discussion: A major contribution of this work is precise identification of multiple concepts related to social connectedness. By integrating annotated examples of both true and false positives, including negations and contextually ambiguous terms, the model better distinguished relevant social context cues from noise. Training on both social isolation and support provided a dual framework for comparative analyses and patient stratification. Conclusion: Transformer-based NLP models, particularly FLAN-T5-Large, demonstrated potential for identifying social isolation and social support in clinical text. These findings support the use of generative AI techniques to enhance detection of social isolation from EHRs, advancing context-aware healthcare analytics.
Qian, L.; Lu, X.; Haris, P.; Yang, Y.
Show abstract
Clinical trials are critical milestones in the drug development pipeline, yet their high failure rates and substantial costs underscore the need for robust predictive models. This study introduces a Heterogeneous Gated Graph Transformer (HGGT) model tailored to predict clinical trial success. Unlike existing methods that typically model trial-related entities in isolation or with homogeneous graphs, HGGT explicitly models the rich heterogeneous relationships among trials, diseases, drugs, genes, targets, abstracts, and eligibility criteria through a gated graph transformer architecture, which dynamically learns and weights multi-type relational interactions to capture complex biological and clinical dependencies. By integrating heterogeneous graph representation with transformer-based context modeling, HGGT effectively captures non-linear, multi-scale interactions across biomedical entities, leading to improved predictive performance for trial success. Experimental results demonstrate that the HGGT model achieves strong performance, with the highest PR-AUC, F1 score, and ROC-AUC across three phases. These findings highlight the potential of graph-based deep learning approaches in optimizing clinical trial design and resource allocation, ultimately accelerating the translation of novel therapies into clinical practice.
Rey-Blanes, A.; Veredas-Morente, J.; Moreno-Barea, F. J.; Veredas, F. J.
Show abstract
Objectives: This study investigates large language models (LLMs) for clinical entity projection across substantial textual transformation. Specifically, we evaluate whether entities annotated in Spanish prostate cancer case reports can be preserved and explicitly projected when the source narratives are transformed into hospital-style clinical progress notes. Entity projection is treated as a generation-driven task, allowing paraphrase, condensation and narrative reorganisation, providing that clinically relevant entities remain recoverable as structured annotations. Methods: A corpus of 109 Spanish prostate cancer case reports was annotated using a silver-standard pipeline combining Spanish biomedical named-entity recognition with rule-based prostate-specific antigen (PSA) and Gleason extractors. The resulting silver-standard annotations were validated on a subset of generated notes against a gold-standard consensus produced by medical experts in prostate cancer. Four LLMs were evaluated for note generation and entity projection: GPT-5.4 Nano, Qwen 3.5:35B-A3B, GLM5 and Claude Sonnet 4.6. Entity-to-Entity (E2E) generation used XML-annotated cases as RAG-supported input, whereas Text-to-Entity (T2E) generation required models to generate and annotate notes directly from plain text cases. Zero-shot and few-shot prompting were tested. Projection quality was measured using precision, recall and F1-score, and complemented by LLM-as-a-judge evaluation using Kimi K2.6. Results: E2E consistently outperformed T2E, indicating that explicit entity-enriched in- put substantially facilitates entity preservation and localisation. GLM5 achieved the best E2E zero-shot result (F1 = 0.915), followed by Claude Sonnet 4.6 (F1 = 0.896). In T2E, few-shot prompting improved performance, with Claude Sonnet 4.6 reaching the highest score (F1 =0.718). Age, Gleason, Disease, Procedure, Duration and negation-related entities were robustly projected, whereas PSA and Dose showed less stable behaviour. Conclusion: LLMs can generate clinically plausible synthetic prostate cancer evolution notes while preserving a substantial proportion of source entities, particularly when explicit semantic annotations are provided as input. However, the lower and more variable performance observed in T2E highlights the difficulty of jointly generating clinical narratives and projecting entities without source-side information, especially for numerical and measure-related entities.
Bin Akter, S.; Akter, S.; Eisenberg, D.; Hill, C.; Lotvola, A.; Fresneda Fernandez, J.; Sarkar Pias, T.; Rafiqul Islam, M.; Islam, H.
Show abstract
Background and Objective: Early and reliable disease prediction from structured clinical data remains challenging when datasets are small, highly imbalanced, and contain limited positive disease cases. Conventional machine learning (ML) and deep learning approaches often struggle to capture clinically meaningful relationships under such low-data representation conditions due to weak statistical associations between features and prediction targets. This study proposes a clinically grounded GPT2-based table-to-text framework for disease prediction using structured healthcare datasets, motivated by the contextual reasoning capability of GPT models to better capture clinically meaningful relationships when statistical learning alone becomes insufficient due to limited data availability. Methods & Materials: Structured clinical records were transformed into physician-style textual descriptions and enriched through GPT4-generated medical paraphrasing to improve minority-class representation while preserving clinical meaning. Both the original and generated clinical texts were used to fine-tune a GPT2 model across four public healthcare datasets, including heart disease, heart failure, chronic kidney disease, and thyroid cancer recurrence. Gradient-based explainable AI analysis was additionally incorporated to identify clinically important features influencing prediction outcomes. Results: The proposed framework demonstrated consistently strong predictive performance with average precision, specificity, sensitivity, and F1-score of 0.96, 0.97, 0.96, and 0.96, respectively. The model achieved improved sensitivity, stronger generalization, and more stable predictive behavior compared with traditional ML, deep learning, transformer-based, and GAN-augmented approaches. Importantly, the framework consistently emphasized clinically meaningful variables even under severe class imbalance where conventional ML and neural network models often struggled. Conclusions: The proposed GPT2-based table-to-text framework provides a practical and clinically interpretable approach for disease prediction from limited structured healthcare data. By integrating contextual clinical reasoning with explainable prediction mechanisms, the framework demonstrates strong potential for early risk detection, transparent clinical decision support, and reliable deployment in real-world low-resource healthcare environments.
Yang, C.-H.; Salvatore, M.; Lu, H.; Zhu, Z.; Tennant, P.; Shi, X.; Ohno-Machado, L.; Khera, R.; Gross, C.; Li, F.; Mukherjee, B.
Show abstract
Electronic health record (EHR)-linked cohorts support association, prediction, and causal studies using longitudinally measured markers of health. However, a lab biomarker measurement is recorded only when a patient first has a medical encounter (visit process) and, a clinician orders the corresponding test and the patient follows through (observation process). These two stages may induce informative presence (IP) and informative observation (IO), respectively. Yet their drivers remain largely uncharacterized, despite evidence that understanding this recording mechanism is essential for selecting appropriate strategies for downstream analysis that treat these markers as longitudinally measured outcomes. We characterize this two-stage recording hierarchy using a stochastic recurrent-event model for the outpatient visit process and a visit-process-weighted generalized estimating equation model for biomarker recording conditional on an outpatient visit. We characterize descriptors of both processes in three EHR-linked cohorts in the US (All of Us [AoU], n=599,423; Yale New Haven Health System [YNHHS], n=319,666; Michigan Genomics Initiative [MGI], n=82,372), reporting descriptive statistics for longitudinal visits and for a panel of 68 lab biomarkers commonly measured in EHRs. We conduct detailed model-based analyses of ten biomarkers spanning multiple domains: routine monitoring, general laboratory assessment, and symptom-triggered testing. These include glucose, hemoglobin A1c [HbA1c], creatinine, hemoglobin [Hgb], white blood cell count [WBC], low-density lipoprotein [LDL] and high-density lipoprotein [HDL] cholesterol, triglycerides, C-reactive protein [CRP], and thyroid-stimulating hormone [TSH]. Across the three cohorts, the median number of outpatient visits ranged from 1.7 to 6.1 per year over a median follow-up of 4.4 to 7.2 years. Among patients with at least one recorded measurement, the median within-person proportion of visits containing a given biomarker ranged from 0.4% to 19.5%, demonstrating that more frequent visits did not necessarily translate into greater per-visit biomarker capture. In the visit-process models, chronic disease burden, and a recent history of outpatient visits were consistently associated with higher visit rates across all three cohorts whereas associations with race, ethnicity, and neighborhood-level income varied across cohorts. In per-visit observation models, the association of covariates depended on the biomarker under consideration; for example, prior cancer diagnosis was associated with more frequent measurement of blood counts but with less frequent measurement of lipids. These findings provide a deeper understanding of how to model who seeks care and what is measured as two distinct recording processes in EHR. Our empirical findings show that the descriptors of these processes vary across cohorts and biomarkers, providing guidance on how to construct these models for downstream longitudinal analyses with irregular EHR visits.
Dhaubhadel, S.; Cohn, J. D.; Bhattacharya, T.; Ribeiro, R. M.; Ganguly, K.; Hengartner, N. W.; Tate, J. P.; Costa, L.; Ho, Y.-L.; Cho, K.; Costa, L.; Beckham, J. C.; Kimbrel, N. A.; Justice, A. C.; McMahon, B. H.
Show abstract
We present a data-driven framework to predict 15-year all-cause mortality using outpatient administrative records for 2.3 million Veterans in the largest integrated U.S. healthcare system. Rather than relying on predefined clinical phenotypes, we used the 1,000 most common outpatient medical codes from each of three data types/modalities: ICD-9 (Dx), Current Procedural Terminology (CPT), and prescription drugs (Rx), encoded as binary features. Using these features, we trained three machine learning (ML) algorithms (logistic regression with lasso, random forest, and a 3-layered feed-forward neural network) to predict 15-year mortality risk. The features were also mapped to variables for the widely used Charlson Comorbidity Index (CCI), Elixhauser, and Veterans Aging Cohort Study (VACS) indices, refitted for 15-year mortality prediction, for baseline comparison. All our models significantly outperformed the widely used CCI, Elixhauser, and VACS indices, with C-statistics ranging from 0.82 to 0.84 versus 0.739 to 0.804 for the baselines. Relative improvements in C-statistics of our approach over the baselines were consistent across different subgroups (age groups of <65 years, those 65+years, Blacks, Hispanics, etc.) Our approach enabled the identification of high-impact predictors with clinical grounding , without requiring hand-curated phenotypes. Cardiovascular diseases and mental health diagnoses/treatments emerged as leading long-term mortality indicators. Using unsupervised ML techniques including PCA and K-means clustering, we associated interpretable patterns and complex interactions between diagnoses and treatments, highlighting comorbidities, disease trajectories, and healthcare utilization patterns. The ability to achieve the predictive performance and algorithmically detect such relationships purely from outpatient data supports the scalability and broad applicability of our framework. This framework not only improves mortality risk stratification over existing clinical indices, but also enables better understanding of how medical codes, regardless of category, interact to predict long-term outcomes.
Chong, J.
Show abstract
We present FHIRBench, a benchmark evaluating six FHIR clinical data serialization strategies across four frontier LLMs (Claude Sonnet 4.5, GPT-5.4, DeepSeek V3.2, Qwen3 32B) on three clinical tasks using 100 stratified synthetic FHIR R4 patient bundles. We employ two evaluation layers: token-level F1 and LLM-as-judge rubric on four clinical dimensions, yielding 7,200 evaluations per layer. Our findings reveal four results. First, serialization significantly impacts quality but the direction diverges between layers: Condensed outperforms Raw JSON on F1 for 3/4 models (Wilcoxon p < 10^-17), while Raw JSON achieves higher judge scores for 3/4 models (p < 10^-7). Narrative achieves 95% of Raw JSON's quality at 83% fewer tokens. Second, model rankings completely reverse between layers -- Claude ranks last on F1 but first on clinical quality (p = 1.0 x 10^-6), demonstrating that single-metric evaluation produces misleading model selection. Third, a significant Model x Serializer interaction (Friedman p = 0.0009) precludes universal format recommendations, with GPT-5.4 favoring Raw JSON while open-weight models favor compressed formats. Fourth, Llama 3.1 70B exhibits 100% inference failure on complex patients despite operating within its nominal context window, revealing a patient-safety gap where AI fails for the patients who need it most. These findings establish that clinical AI systems require model-aware serialization middleware, multi-layer evaluation frameworks, and capacity verification before deployment. Code and data publicly available.
Gao, Y.; Cui, Y.
Show abstract
Large-scale clinical and biomedical datasets increasingly contain both diverse subgroup attributes (e.g., demographic or clinical subgroups) and multiple prediction targets. Although various machine learning approaches can address subgroup differences or multi-target prediction, they often consider these aspects independently rather than jointly. To more effectively capture the shared and subgroup-specific information in such complex datasets, we propose the Integrative Transfer Network (ITN), a deep neural network designed to leverage data across subgroups and multiple related outcomes simultaneously. In extensive experiments, including time-to-event and classification tasks where demographic subgroups and multiple disease end-points are prevalent, ITN demonstrates consistent improvements in subgroup-specific prediction by borrowing strength from other subgroups and outcomes. We envision ITN as a unified frame-work for learning from heterogeneous datasets where subgroup-specific insights are critical.
Badhon, S. M. S. I.; Adibuzzaman, M.; Mosa, A. S. M.; Bozdag, S.; Cleveland, A. D.; Ding, J.; Hossain, K. S. M. T.
Show abstract
Objective: Acute kidney injury (AKI) affects a large proportion of patients in the intensive care unit (ICU) and is a major contributor to morbidity, mortality, and cost. Although electronic health records (EHRs) capture rich longitudinal data, many predictive models fail to detect AKI early enough for effective intervention. Non-temporal methods such as logistic regression and XGBoost treat patient history as aggregated risk factors, discarding the temporal evolution of clinical state. A recent trend is to employ temporal models, such as recurrent neural networks, to capture sequential patterns, but these models struggle with irregular sampling and limited long- range contextual awareness. To address the challenge, we propose RenalTransLSTM, a hybrid temporal deep learning framework for early, multi-horizon AKI prediction and identification of modifiable risk factors. Methods: RenalTransLSTM integrates Long Short-Term Memory (LSTM) networks with Transformer encoders to model both local temporal dynamics and global contextual depen- dencies in ICU time-series data. Using 48-hour patient histories from MIMIC-IV (61,735 admissions), the model predicts AKI at 6-, 12-, and 24-hour lead times. We benchmark the model against SVM, XGBoost, LSTM, TG-LSTM, and a Transformer, and apply Integrated Gradients and counterfactual analysis to identify modifiable risk factors. Results: RenalTransLSTM outperforms all baselines across most horizons and metrics, achiev- ing AUROC above 0.90 and F1-scores reaching 0.85 while maintaining balanced precision and recall on imbalanced data. Ablation studies confirm that combining LSTM and Transformer components improved robustness and predictive performance. Counterfactual analysis identifies clinically meaningful, modifiable risk factors associated with AKI progression. Conclusion: RenalTransLSTM offers an effective, interpretable framework for early AKI prediction in the ICU, supporting proactive intervention and clinical decision support.
DU, J.; Deng, G.
Show abstract
While Directed Acyclic Graphs (DAGs) are essential for causal inference, their construction often relies on expert heuristics, which bypasses systematic evidence synthesis and creates a critical "evidence retrieval gap" in causal modeling. This study introduces EpiKG2DAG, a framework that supports evidence-anchored candidate DAG generation by transforming unstructured biomedical abstracts into structured epidemiological associations. We utilized DeepSeek-V3 to extract exposure-outcome association triplets from 189,266 abstracts and employed SapBERT for semantic normalization against UMLS concepts. The resulting Epidemiological Knowledge Graph (EpiKG) enables the automated identification of candidate confounders, mediators, and colliders based on graph-theoretic motifs and literature-derived evidence. A case study on COVID-19 and AKI demonstrates that the framework uncovers non-obvious confounders, such as air pollution, while ensuring evidence traceability. This work contributes to the field by mitigating the knowledge-acquisition bottleneck and providing a transparent, reproducible foundation for evidence-based causal modeling.
Song, Q.; Ni, C.; Liu, W.; Li, Y.; Malin, B. A.; Yin, Z.
Show abstract
Automatic coding from clinical notes has been studied extensively for International Classification of Diseases (ICD) codes, yet broad Current Procedural Terminology (CPT) and Healthcare Common Procedure Coding System (HCPCS) recommendation remains comparatively underexplored. Existing studies often focus on one specialty, a limited code vocabulary, or a single model family, leaving it unclear how different artificial intelligence (AI) paradigms perform under a common, clinically meaningful evaluation. We formulate CPT and HCPCS coding as an AI-assisted recommendation task in which a physician or professional coder reviews a short, ranked list of candidate codes supported by the clinical note. Using operative notes from Vanderbilt University Medical Center (VUMC) and discharge summaries from Medical Information Mart for Intensive Care IV (MIMIC-IV), we compare lexical retrieval, Clinical-Longformer, GPT-5.6-Sol, MedGemma-27B, and an inspectable agentic-style retrieve-and-verify system under a controlled review budget. Micro-averaged recall within a fixed number of recommendations measures whether reference codes reach the reviewable list; micro-F1 is reported only where reference labels are sufficiently complete. Zero-shot GPT-5.6-Sol achieves the highest recall within five and ten candidates: 0.717 and 0.800 on VUMC and lower-bound values of 0.689 and 0.738 on MIMIC-IV. The retrieve-and-verify system reaches 0.695 and 0.784 on VUMC and lower-bound values of 0.575 and 0.657 on MIMIC-IV, with a candidate-linked evidence window attached to each retained recommendation. Diagnostic analyses reveal distinct failure sources, including output-length underfilling, confusion among closely related codes, out-of-knowledge-base generation, and incomplete evidence support. These findings establish a systematic evaluation framework for procedure-code recommendation and identify practical requirements for future systems that are accurate, review-efficient, and grounded in clinical evidence.
Oehring, D.
Show abstract
Background Averagebased summaries serve individual patients poorly PORTRAIT is a calibrated abstentionaware tool that describes where one patient sits relative to a reference population across 12 cardiometabolic markers how confident that placement is and which features drive it PORTRAIT describes it does not diagnose or predict Abstention is a designed feature given the known limits of conditional coverage Methods Conformal calibration was combined with distributionfree coverage bounds quantileregression coordinates and copulabased joint structure A frozen reference cohort n9421 supplied fixed calibration a heldout cohort n2247 tested transportability across six strata A release gate required the minimum perslice coverage to hold across 4 of 5 seeds Coverage was retested under survey weighting to the US adult population Coherence was reported as a descriptive joint coordinate Discrimination was summarised with Harrells C and multiplicity controlled by BHFDR Interface conformance was assessed against defined requirements Nielsen heuristics and WCAG 22 AA with attention to automation bias and riskgraph design Results The frozen reference held all six strata within band 08640903 at abstention 0113 whereas a resplit undercovered to 071 at abstention 0227 coverage survived survey weighting The release gate passed on 4 of 5 seeds at abstention 0101 against a nominal 090 and in the frozenreference configuration that ships all six strata held inside the calibrated band 08640903 Coherence showed orthogonality 0444 to raw extremity and correlated 0892 with a copulaMahalanobis distance while remaining deliberately nonidentical so it adds perfeature information Two transfer tests returned negatives the ocular transfer did not hold coverage at thinn Adding coherence changed mortality discrimination by deltaC 00047 Interface requirements moved from 142718 to 38147 METPARTIALUNMET Nielsen severity resolved 7 of 10 issues WCAG 22 AA text criteria passed Conclusions PORTRAIT situates a patient against a frozen reference holds coverage under survey weighting to the US adult population and abstains when calibration cannot be supported The headline result is that the frozen reference held coverage where a resplit did not
Wang, Y.; Stroh, J. N.; Ghosh, D.; Sirlanci, M.; Hripcsak, G.; Bennett, T. D.; Albers, D.
Show abstract
Clinical decisions for determining optimal patient-specific interventions are complicated prediction tasks that rely on health care professionals' understanding of physiological mechanisms and their dynamics. These decisions are challenged by (a) observational data sparsity and (b) patient heterogeneity. Here, we focus on estimating and forecasting specific physiological properties--that are not explicitly present in clinical observations--to provide additional features using only data available bedside at the time of decision-making. Mechanistic models of physiological system(s), e.g., physiological ordinary differential equation (ODE) models, provide pathways to compensate for data sparsity by synchronizing the model with observations of an individual patient using data assimilation (DA). However, DA used in a standard computational workflow to estimate constant model parameters from presently-known data is less effective at optimizing state forecasts of the model governed by physiological processes that evolve before new observations are available. Stated simply, we cannot forecast the future evolution of the model because we cannot forecast model parameters. To support next-generation clinical decision support, we develop a new DA and machine learning (ML) hybrid pipeline to estimate and forecast individual future physiological processes by forecasting ODE model parameters. This pipeline overcomes model and DA workflow limitations by stacking a DA-estimated posterior empirical distribution of physiological parameters with longitudinal ML forecasting models. We work within the context of glycemic management in an ICU using EHR data to construct and test a use case. We use synthetic data and real-world clinical data to validate the integrated pipeline and quantify uncertainties.
Bukhari, S. A. C.; Hayder, N. S.; Wajahat, I.
Show abstract
Existing evaluations of healthcare AI often treat interoperability as a technical infrastructure issue rather than a factor that directly influences the safety and reliability of clinical AI systems. Yet the quality of Fast Healthcare Interoperability Resources (FHIR) implementation affects whether AI models can operate accurately, fairly, securely, and effectively in real clinical settings. We present FHIRTrustBench, a benchmark for assessing the readiness of FHIR-based clinical AI systems across five complementary dimensions: FHIR implementation quality, AI validation, clinical workflow integration, trustworthiness assessment, and governance readiness. Each dimension is mapped to a distinct category of downstream deployment failure risk. We applied FHIRTrustBench to a corpus of 10 representative sources spanning interoperability standards, implementation studies, electronic health record integration research, healthcare large language model research, and governance frameworks. Each source was scored individually and traceably against the five-dimension rubric. FHIR Specificity achieved the highest dimension mean at 1.3 out of 2.0, while AI Validation received the lowest at 0.3. Even category-leading sources that scored a maximum 2.0 on FHIR Specificity scored 0 on AI Validation. Prospective external validation was reported in no source, and Governance Readiness remained at or below 1.0 across every category. We further identify five interoperability-related AI failure pathways, spanning data integrity, semantic consistency, security, clinical workflow, and generative AI grounding, and propose a deployment lifecycle framework and reporting checklist that translate benchmark scores into deployment-readiness decisions for developers, healthcare organizations, and regulators. FHIRTrustBench provides a practical and reproducible basis for assessing FHIR-enabled clinical AI before deployment and can evolve as interoperability standards and clinical evidence mature.
Santos, R. d. P.; Tinoco Patricio, A. d. O.; Gama, P. H.; Freitas, L. M. D.; Ribeiro, K. R.
Show abstract
Objective: To construct and evaluate, in an exploratory manner, a pathophysiologic rationale link- ing biological pathways derived from the peripheral transcriptome in ischemic stroke (IS) to nursing diagnoses in the NANDA-I 2024-2026 taxonomy, while emphasizing that this association is not di- rect, deterministic, or automatically inferable from textual similarity with large language models (LLMs). Methods: A computational study was conducted using public secondary data from the Gene Ex- pression Omnibus series GSE16561, which includes 63 peripheral blood samples: 39 from indi- viduals with IS and 24 from healthy controls. The pipeline integrated transcriptomic analysis and functional enrichment, semantic mapping through ClinicalBERT embeddings, and mechanistic and clinical-conceptual judgment using Claude Sonnet 4.6 as a judge. The judgment stage was treated as the central interpretive layer, designed to mediate the transcriptome, pathophysiology, functional manifestation, and NANDA-I diagnosis. Results: The analysis identified a bimodal transcriptomic pattern, with activation of pathways re- lated to innate immunity and suppression of pathways related to adaptive immunity. Semantic map- ping generated 158 pathway-diagnosis pairs. The Spearman correlation between cosine similarity and the mechanistic score was negative and statistically significant (rho = -0.243; p = 2.09e-03), but weak in magnitude. This effect size indicates that semantic similarity explained less than 6% of the variance in mechanistic plausibility, reinforcing the insufficiency of embeddings as a stand- alone criterion. Of the 158 pairs, 14 were classified as high concordance, 8 as moderate, and 136 as divergent. Conclusion: The main value of this study lies in demonstrating that translating biological pathways into nursing diagnoses requires pathophysiologic, functional, and clinical-conceptual mediation. The prioritized pairs represent mechanistically plausible hypotheses for future research, without implying causality, direct clinical confirmation, or immediate care recommendations.
Wang, N.; Kakadiaris, A.; Li, C.; Wang, R.; Ahn, J.; Wang, Y.; Fu, S.
Show abstract
Symbolic clinical natural language processing (NLP) systems remain widely used for extracting clinical concepts from electronic health record (EHR) narratives, but maintaining rule resources requires extensive manual error analysis and rule refinement. This study investigates whether large language models (LLMs) can assist in identifying extraction errors and generating candidate rules to improve symbolic clinical NLP systems. Using error reports derived from a multi-site evaluation of a previously validated symbolic model for cognitive and neuropsychiatric-related clinical concepts, we developed a human-in-the-loop framework, REFINE. The framework first uses LLMs to classify extraction errors and generate explanatory reasoning, which can then be incorporated into prompts for rule generation. Three LLMs (GPT-5.2, GPT-4o, GPT-4o-mini) were evaluated under four prompting conditions. LLM-generated rule sets improved performance compared with the baseline NLP-CAM system, increasing F1-score from 0.37 to 0.58. These findings suggest that LLMs can support scalable rule refinement for symbolic clinical NLP systems.
Xue, X.; Frydman-Gani, C.; Arias, A.; Perez Vallejo, M.; Londono Martinez, J. D.; Valencia-Echeverry, J.; Castano, M.; Freimer, N. B.; Lopez-Jaramillo, C.; Olde Loohuis, L. M.
Show abstract
Background: Free-text notes in electronic health records (EHRs) contain fine-grained psychiatric information that is essential for psychiatric research and clinical care, and often absent or under-recorded in structured codes alone. Clinical natural language processing (cNLP) can support extraction of this information from EHR notes, yet Spanish-language cNLP remains under-developed. Moreover, broad evaluations comparing multiple encoder-based language models across extensive, fine-grained psychiatric concept sets remain scarce, and it remains unclear how these models compare with traditional NLP (tNLP) systems and much larger generative large language models (LLMs). In addition, cross-site performance of fine-tuned models is rarely tested, and limited annotated training data remains a major challenge, especially for rare symptoms. Objectives: We aimed to advance scalable, global psychiatric cNLP by fine-tuning multiple encoder-based models with differing architectures and pre-training strategies for detecting fine-grained psychiatric concepts in Spanish EHRs. We further evaluated the impact of augmenting the fine-tuning data with precision-weighted weak labels for less-frequent concepts, and compared the performance of the encoder-based models to that of tNLP and a fine-tuned generative LLM trained on the same data. Finally, we evaluated model cross-site generalizability on an external EHR dataset. Methods: Three encoder-based models (BETO, XLM-RoBERTa-large, and bsc-bio-ehr-es) were fine-tuned on 1,642 clinician-annotated EHR documents from Colombia to detect 110 psychiatric concepts in Spanish text. To address the limited annotated examples available for less-frequent concepts, 12,000 additional documents were weakly-labeled for less-frequent concepts using tNLP, and incorporated into the fine-tuning data with labels weighted by pattern precision. Models were compared with tNLP and a generative LLM, and evaluated on an external EHR dataset from another psychiatric hospital in Colombia. Results: Encoder model performance varied substantially, with macro-F1 ranging from 0.64 to 0.81. BETO achieved the highest macro-F1 (0.81; median F1=0.88 [IQR=0.77-0.96]). Adding precision-weighted weak labels for less-frequent concepts improved BETO's overall macro-F1 to 0.83 and increased mean F1 for the 55 augmented concepts from 0.82 to 0.86. Under matched fine-tuning conditions, fine-tuned BETO and the tNLP method were equivalent in F1, whereas the LLM significantly outperformed BETO in F1. After weak-label augmentation, BETO significantly outperformed tNLP in F1 (PFDR<.001) and narrowed the performance gap with the LLM, although equivalence was not established. Lastly, fine-tuned BETO maintained reasonably strong performance on data from an external hospital not used for model fine-tuning (out-of-domain macro-F1=0.78). Conclusions: General-purpose pre-trained encoders had strong performance for psychiatric concept extraction from Spanish EHRs. Weak-label augmentation improved BETO's performance and strengthened results relative to a tNLP baseline, while reducing, but not eliminating, the performance gap with a much larger fine-tuned generative LLM. These findings highlight the utility of these relatively lightweight models for scalable, accurate and reproducible detection of psychiatric concepts in Spanish-language EHRs.